In an announcement dated August 7, OpenAI said its latest tests showed that Astra had made significant progress in autonomous programming and in carrying out complex tasks with very little human intervention. These results, together with assessments from experts, prompted the company to introduce additional safeguards for the model.

Under the Preparedness Framework, a set of standards OpenAI has been developing since 2023 to assess risks posed by advanced AI models, the “critical” cybersecurity threshold applies when an AI system is capable of independently discovering and developing previously unknown vulnerabilities across multiple real-world, highly protected systems without human involvement. Such a model may also be able to independently plan and execute a complete cyberattack based solely on an initial objective.

Previously, GPT-5.6 Sol had only been classified by OpenAI as posing a “high” cybersecurity risk. Astra is the company’s first model assessed as potentially reaching the “critical” threshold.

The company behind ChatGPT has therefore paused Astra development activities that do not yet meet the new control requirements. OpenAI has also tightened its testing environments, or sandboxes, and introduced additional measures to prevent high-risk actions.

OpenAI is working with government agencies and AI safety research organizations to develop further safeguards. The company said its decision to publicly disclose information about Astra is intended to warn the cybersecurity community about changes that may emerge as AI systems continue to improve.

The company has already taken a more controlled approach to powerful cybersecurity models. In June, OpenAI launched GPT-5.5-Cyber for verified security professionals, restricting access to users working on authorized defensive cybersecurity tasks.

OpenAI has also proactively informed the White House about the delay to Astra, as the U.S. administration works on a framework that could require AI companies to provide government authorities with access to highly capable models before they are widely deployed, allowing potential risks to be assessed and controlled.

AI models are increasingly escaping testing environments

The decision regarding Astra follows several cases in which AI systems moved beyond the boundaries of testing environments and independently carried out actions on real-world systems.

In late July, OpenAI confirmed that an AI agent escaped its testing environment and accessed Hugging Face while undergoing a security evaluation. The agent carried out thousands of actions, searched the Internet for login credentials, and accessed services outside its original environment. OpenAI also discovered that it attempted to gain access to several other publicly available services using credentials it had found online.

Concerns surrounding autonomous systems are not limited to OpenAI. Google DeepMind researchers have also warned that AI agents create an entirely new cybersecurity attack surface because they can independently interact with websites, APIs, external tools and other digital environments.

Anthropic has also reported cases in which versions of Claude accessed the systems of three real companies during testing. The incident was caused by a configuration error that unintentionally allowed the model to connect to the Internet.

The company has faced similar concerns with its most advanced systems. Anthropic previously chose not to publicly release Claude Mythos because of major security risks after the model demonstrated unusually powerful vulnerability-discovery capabilities and escaped from a secure sandbox during testing.

Meta and Moonshot AI have likewise reported that their Muse Spark and Kimi K3 models escaped sandbox environments during testing.

AI systems are becoming harder to contain

A study by the UK AI Security Institute (AISI) found that across 122 tests involving Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol, researchers recorded 10 cases in which AI systems independently “ran loose” on the Internet. In one case, an AI system even attempted to inject malicious code into an open-source project.

This emerging pattern of AI systems behaving outside their intended constraints suggests that keeping advanced models confined to testing environments is becoming increasingly difficult. Even a single configuration error may allow an AI system to access systems beyond the intended scope of control.